Papers with policy network

8 papers
Adaptive Multi-pass Decoder for Neural Machine Translation (D18-1)

Copied to clipboard

Challenge: End-to-end neural machine translation (NMT) has attracted increasing attention in recent years.
Approach: They propose an adaptive multi-pass decoder which introduces a flexible multi- pass polishing mechanism to extend the capacity of NMT via reinforcement learning.
Outcome: The proposed architecture improves Chinese-English translation with 1.55 BLEU . the proposed architecture adopts a flexible multi-pass polishing mechanism .
Actor-Double-Critic: Incorporating Model-Based Critic for Task-Oriented Dialogue Systems (2020.findings-emnlp)

Copied to clipboard

Challenge: In order to improve the sample-efficiency of deep reinforcement learning, we implemented imagination augmented agent (I2A) in spoken dialogue systems (SDS).
Approach: They propose to use an actor-double-critic to improve the stability and overall performance of imagination augmented agent (I2A) in spoken dialogue systems.
Outcome: The proposed model-based agent (ADC) improves the stability and sample-efficiency of deep reinforcement learning (DRL) on a restaurant booking task.
Struct-XLM: A Structure Discovery Multilingual Language Model for Enhancing Cross-lingual Transfer through Reinforcement Learning (2023.emnlp-main)

Copied to clipboard

Challenge: Existing methods require syntactic labels that are difficult to obtain and of poor quality for low-resource languages.
Approach: They propose a syntactic alignment model that leverages reinforcement learning to discover universal syntaktic structures for cross-lingual PLM alignment.
Outcome: The proposed model improves cross-lingual representation alignment on the XTREME benchmark.
End-to-End Reinforcement Learning for Automatic Taxonomy Induction (P18-1)

Copied to clipboard

Challenge: Existing methods for automating taxonomy induction often divide the problem into two subtasks . a novel end-to-end reinforcement learning approach is proposed to improve the accuracy of such methods.
Approach: They propose an end-to-end reinforcement learning approach to automatic taxonomy induction from a set of terms.
Outcome: The proposed approach outperforms state-of-the-art methods on two public datasets of different domains.
Inverse Reinforcement Learning for Text Summarization (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing studies show that inverse reinforcement learning (RL) training has certain disadvantages such as object mismatch and exposure bias.
Approach: They propose inverse reinforcement learning (IRL) as an effective paradigm for training abstractive summarization models.
Outcome: The proposed model outperforms MLE and RL baselines on ROUGE, coverage, novelty, compression ratio, factuality, and human evaluations.
Multi-Label Few-Shot Learning for Aspect Category Detection (2021.acl-long)

Copied to clipboard

Challenge: Existing few-shot learning methods focus on single-label predictions, which can not work well for ACD since a sentence may contain multiple aspect categories.
Approach: They propose a few-shot learning method that uses the prototypical network to learn aspects from a set of aspects.
Outcome: The proposed method significantly outperforms baseline methods on three datasets.
FLAG-TRADER: Fusion LLM-Agent with Gradient-based Reinforcement Learning for Financial Trading (2025.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) have impressive reasoning capabilities in financial tasks, but struggle with multi-step, goal-oriented scenarios in interactive financial markets.
Approach: They propose a framework that integrates large language models with gradient-driven reinforcement learning (RL) policy optimization.
Outcome: The proposed framework improves performance in trading and other financial domain tasks.
LeLoRA: Learnable Low-Rank Adaptation of Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches to fine-tuning large language models (LLMs) rely on manually specified and fixed hyperparameters, resulting in suboptimal performance and low parameter efficiency.
Approach: They propose a framework that allows for dynamically learned adaptive adaptation strategies to be used to fine-tune large language models.
Outcome: The proposed framework outperforms baselines in adapting large language models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations